Papers with evaluation error
Poller: Are LLMs Suitable for Evaluating Poetry Understanding Task? (2026.findings-acl)
Copied to clipboard
| Challenge: | Traditional methods for poetry evaluation are expensive and unsuitable for large-scale data. |
| Approach: | They propose a method leveraging Large Language Models to evaluate poetry understanding tasks using Large Language models. |
| Outcome: | The proposed method reduces the evaluation error between LLMs and humans by adopting the poet's perspective. |
SciCompanion: Graph-Grounded Reasoning for Structured Evaluation of Scientific Arguments (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing retrieval-augmented generation (RAG) methods fail to provide deep, relational understanding of scientific literature. |
| Approach: | They propose a graph-grounded reasoning framework for structured scientific evaluation that uses multi-hop reasoning to iteratively construct contextual graphs and generate structured critiques. |
| Outcome: | The proposed framework reduces evaluation error by over 30% compared to baselines and allows smaller models to outperform larger models. |